Skip to content

[Core] Scope PCP-DP validation to GPU manager - #54523

Merged
vllm-bot merged 6 commits into
vllm-project:mainfrom
pisceskkk:codex/pcp-dp-validation-scope
Sep 8, 2026
Merged

vllm-bot merged 6 commits into
vllm-project:mainfrom
pisceskkk:codex/pcp-dp-validation-scope

Conversation

@pisceskkk

@pisceskkk pisceskkk commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Move the current PCP-with-DP capability check from the platform-independent ParallelConfig validator to the GPU MRV2 PCPManager. This keeps the existing GPU behavior while allowing out-of-tree hardware backends to provide their own PCP+DP support.

Test Plan

  • Run Ruff lint and formatting checks on the changed files.
  • Run the standard PR CI checks.

Test Result

  • Ruff lint: passed.
  • Ruff format check: passed.
  • Python compileall: passed.

Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR is described.
  • The test plan is included.
  • The available test results are included.
  • No documentation update is required for this validation-scope change.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-08-31T08:04:17.004970Z acd364b PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
@pisceskkk
pisceskkk force-pushed the codex/pcp-dp-validation-scope branch from acd364b to 29ce06a Compare August 31, 2026 08:02

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: acd364b85a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vllm/v1/worker/gpu/pcp_manager.py Outdated
@mergify mergify Bot added the mrv2 Model Runner V2 specific label Aug 31, 2026
@ZJY0516 ZJY0516 added the ready ONLY add when PR is ready to merge/full CI is needed label Aug 31, 2026
@ZJY0516

ZJY0516 commented Aug 31, 2026

Copy link
Copy Markdown
Member

/ci run

@ZJY0516
ZJY0516 enabled auto-merge (squash) August 31, 2026 08:45
@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #86328 for commit 29ce06ab9823.

@github-actions

Copy link
Copy Markdown

✅ @pisceskkk, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@ZJY0516
ZJY0516 disabled auto-merge August 31, 2026 08:50
@ZJY0516

ZJY0516 commented Aug 31, 2026

Copy link
Copy Markdown
Member

/ci cancel

Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
@mergify mergify Bot removed the needs-rebase label Sep 2, 2026
@pisceskkk

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #86747.

@pisceskkk

Copy link
Copy Markdown
Contributor Author

/amd-ci retry

@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

✅ Queued 1 failed job(s) for retry in Buildkite AMD CI #12538.

@coderabbitai

coderabbitai Bot commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Team

Run ID: 995c6eb9-0b1f-4d3e-8b09-1d119773dcca

📥 Commits

Reviewing files that changed from the base of the PR and between 9cc7793 and efb1d92.

📒 Files selected for processing (3)
  • vllm/config/parallel.py
  • vllm/platforms/cuda.py
  • vllm/platforms/rocm.py
💤 Files with no reviewable changes (1)
  • vllm/config/parallel.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • vllm/platforms/rocm.py
  • vllm/platforms/cuda.py

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes
    • Improved validation for parallel execution configurations.
    • CUDA and ROCm now clearly reject unsupported combinations of prefill context parallelism and data parallelism.
    • Configurations on other platforms can continue to the appropriate existing validation checks, allowing supported combinations to proceed correctly.

Walkthrough

The parallel configuration validator now permits PCP with data parallelism. CUDA and ROCm platform validators reject this combination with platform-specific ValueError checks.

Changes

PCP and data parallelism validation

Layer / File(s) Summary
Allow combined parallel configuration
vllm/config/parallel.py
The validator no longer rejects configurations that combine PCP and data parallelism.
Enforce platform restrictions
vllm/platforms/cuda.py, vllm/platforms/rocm.py
CUDA and ROCm checks now reject configurations where both PCP and data parallelism exceed 1.

Estimated code review effort: 2 (Simple) | ~10 minutes

Merge Risk: ⚪ Minimal · up to efb1d

PCP and data-parallelism validation is now enforced by CUDA and ROCm platform checks while allowing other backends to define their own support. No current merge-blocking risk is identified.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 2 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: PCP-DP validation is scoped to the GPU manager or platform layer.
Description check ✅ Passed The description accurately explains the validation move, preserved GPU behavior, support for out-of-tree backends, and test results.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@pisceskkk

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

✅ The previous CI build is still running: https://buildkite.com/vllm/ci/builds/86747

@pisceskkk

Copy link
Copy Markdown
Contributor Author

/amd-ci retry

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

✅ Queued 3 failed job(s) for retry in Buildkite AMD CI #12631.

@coderabbitai

coderabbitai Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@vllm-bot
vllm-bot merged commit 7c2f1ff into vllm-project:main Sep 8, 2026
8 of 9 checks passed
@github-project-automation github-project-automation Bot moved this from Todo to Done in AMD Sep 8, 2026
@github-project-automation github-project-automation Bot moved this from Ready to Done in NVIDIA Sep 8, 2026
wzx0726 added a commit to wzx0726/vllm-ascend that referenced this pull request Sep 9, 2026
Rely on vllm-project/vllm#54523 to scope PCP+DP rejection to CUDA and
ROCm. Remove the Ascend validator wrapper, cached schema rebuilding,
dedicated tests and patch documentation. Keep dummy execution handling.

Validation: targeted Ruff checks, syntax checks and git diff --check.
Runtime compatibility with vLLM 7c2f1ff remains pending.

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
weijinqian0 pushed a commit to vllm-project/vllm-ascend that referenced this pull request Sep 9, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
ItsRoy69 pushed a commit to ItsRoy69/vllm that referenced this pull request Sep 10, 2026
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Signed-off-by: Jyotirmoy Roy <jyotirmoyroy649@gmail.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: tianming2009 <13246728590@163.com>
LucasWilkinson added a commit to LucasWilkinson/vllm that referenced this pull request Sep 14, 2026
vllm-project#54523 moved the ParallelConfig check into CudaPlatform/RocmPlatform
check_and_update_config after this series branched; the series removes the
ParallelConfig copy, so remove the platform copies too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

Add Ascend MRV2 handling for PCP combined with DP, reusing the upstream
PCP dispatch sizing and input partitioning paths.

When a DP replica becomes idle, its dummy batch bypasses PCP
partitioning. Saved PCP attention state may be absent or belong to the
preceding real batch. Runtime dummy inputs use runner buffers, while
decode graphs capture persistent PCP buffers.

- Build dummy PCP attention contexts from the current batch, block
tables and rank-local slot mappings, including the DSA metadata path.
- Refresh persistent PCP buffers before dummy execution. Keep the
implementation in `AscendPCPManager`;
`NPUModelRunner.prepare_dummy_attn` delegates or takes the existing
non-PCP path.
- Synchronize replicated speculative drafts independently of the
target's PCP-local DP state, with every DP replica participating,
including idle and decode replicas.
- Restrict PCP+DP graph configuration to eager or `FULL_DECODE_ONLY`.

Upstream validation dependency:
vllm-project/vllm#54523, merged as
[7c2f1ff4958eaf0818405e9192c71608fe4a16b1](vllm-project/vllm@7c2f1ff),
moves the PCP+DP restriction from common `ParallelConfig` validation to
CUDA/ROCm platform checks. This PR relies on that change and contains no
Ascend validator bypass or Pydantic schema rebuilding.

Related upstream PR: vllm-project/vllm#53867
([Feature][PCP] Support decode-only FULL CUDA graphs).

Once the paired vLLM includes #53867, adapt its PCP
`prepare_inputs_to_capture` entry point to create `AscendInputBatch`
directly in persistent PCP buffers. After capture and real-to-idle
replay validation, remove `AscendPCPManager.prepare_dummy_attn` and the
runner override, as tracked by `TODO(wzx0726)`. The Ascend dummy
attention-context handling and replicated-draft DP synchronization
remain necessary.

### Does this PR introduce _any_ user-facing change?

Adds Ascend-side MRV2 handling for PCP combined with DP. No new CLI
options or environment variables are introduced.

Rebased onto Ascend main `ad86348b0cb2324d643df6a496f2d9c3879e4481`. The
draft conflict resolution preserves vLLM 0.28.0 `num_tokens_across_dp`
forwarding and applies fresh synchronization for replicated drafts on
the main2main `dp_sync` interface.

**Compatibility and remaining validation:**

- The previously tested baseline was Ascend
`b1c91857c16bde10c2fa7f5d7548b7666a86d4bb` with vLLM
`e6bfe03ad73a3330cb427885aa90d97a12e1c704`.
- Removing the validator bypass requires vLLM #54523. Ascend main
currently pins `b2f685834a6456197e7033966fdef52a23f1abcd`, which
predates that change and still rejects PCP+DP. The local vLLM checkout
remains at the user-selected `7c2f1ff4958eaf0818405e9192c71608fe4a16b1`.
This PR does not change the main2main pin or reintroduce the bypass.
- Compatibility work against that newer revision is still pending,
including the added `prepare_dummy_attn` argument and removal of
`InputBatch.max_seq_len_np`. The current branch is not yet validated to
run against that revision.
- Multi-card Attention/MLA/DSA inference, real-to-idle DP transitions
during `FULL_DECODE_ONLY` replay, speculative-decoding numerical
correctness, and EP/TP combinations require model-level validation on
the final paired sources.

### How was this patch tested?

The latest rebase onto `ad86348b0` resolves an append-location conflict
in `test_pcp_manager_v2.py`, retaining both the main DCP padding
regression and this PR's tests. AST comparison confirms that all main
tests and the final PR tests are preserved, and the other changed files
match the automatic merge. Ruff lint/format, syntax parsing, and `git
diff --check` passed. Runtime tests were not rerun for this
conflict-only update.

For the preceding rebase onto `f8481287a`, **6 targeted candidate-method
cases passed** in an isolated process in the Ascend container: five
main2main draft-sync cases and one mocked v0.28.0 argument-routing case.
The exact candidate `propose` method was loaded into the installed
runtime classes; native dependencies and the rest of the runtime
remained at their installed versions. This is method-level regression
evidence, not full rebased-source or dual-version runtime validation.

All eight changed Python files passed Ruff lint/format checks, syntax
parsing, and `git diff --check`. Full unit-suite and model validation on
the final paired sources remain pending.

The runtime adaptation suite previously passed on the earlier baseline
in an isolated Ascend-container test directory, using CPU tensors and
mocks for device kernels:

```bash
python3 -m pytest -q tests/ut/worker/test_pcp_manager_v2.py tests/ut/worker/test_model_runner_v2.py tests/ut/worker/test_attn_utils_v2.py tests/ut/worker/test_mtp_pcp_speculator_v2.py
```

- **75 passed** (14 existing `torch.jit` deprecation warnings). This is
prior-baseline evidence, not a test result for vLLM `7c2f1ff`.
- Covers missing/stale dummy PCP state, rank-local mappings, persistent
buffer contents and storage, DSA metadata forwarding, non-PCP fallback,
graph restrictions, and replicated draft DP synchronization.
- The removed validator bypass's dedicated tests are also removed; the
existing unrelated platform test is retained.
- For the removal, targeted Ruff lint/format checks, Python syntax
checks and `git diff --check` passed. No runtime tests were rerun
against the upgraded vLLM. The local `format.sh ci` entry point remains
unavailable because its shell lacks `pre-commit`.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wzx0726 <zhexuanwu12@gmail.com>
Signed-off-by: like-0517 <ithwlike@126.com>
LucasWilkinson added a commit to LucasWilkinson/vllm that referenced this pull request Sep 16, 2026
vllm-project#54523 moved the ParallelConfig check into CudaPlatform/RocmPlatform
check_and_update_config after this series branched; the series removes the
ParallelConfig copy, so remove the platform copies too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Tflowers-0129 pushed a commit to vllm-project/vllm-ascend that referenced this pull request Sep 23, 2026
### What this PR does / why we need it?

#### Change Summary

The PR advances the main2main lane to vLLM v0.30.0 (commit
`ced6857afa0ea7b2e3f0846a62e1394e90f15607`), adapting vllm-ascend to
every upstream change in the `84030bbe` -> `4991f97` -> `ced6857` range.
Because both CI lanes now install vLLM v0.30.0, all
`vllm_version_is("0.29.0")` forks are permanently false and are
collapsed to the main behavior.

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash `84030bbe` -> `4991f97` -> `ced6857` (v0.30.0 tag commit) |
| `.github/vllm-release-tag.commit` | — | Bumped the release boundary to
`v0.30.0` |
| `.github/workflows/pr_test.yaml` | — | Added a dual-version cpu-ut
matrix (`vllm_versions`) for the main2main lane; temporarily commented
out pre-commit/mypy and set `cpu-ut` to `if: false`; dropped the
`needs.cpu-ut` requirement from the ready gate; added a
main2main-specific cpu-ut failure hint |
| `Dockerfile` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `README.md` | — | CI notes updated to v0.30.0 |
| `README.zh.md` | — | CI notes updated to v0.30.0 |
|
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_num_nans.py`
| — | `VLLM_VERSION` default 0.29.0 -> 0.30.0 |
| `tests/ut/_310p/test_model_runner_v2_310p.py` | v0.30.0 boundary |
Dropped the `vllm_version_is` gate test; `_needs_kv_cache_zeroing_310p`
always uses `spec_config.use_eagle_block_drop()` |
| `tests/ut/core/test_dyntra_lb_scheduler.py` | — | Removed the 0.29
`KVConnectorBlockState.block_ids` assertions; assert the `req_ids` form
only |
| `tests/ut/core/test_scheduler_connector_block_state.py` | — | Removed
the 0.29 `block_ids` snapshot branch |
| `tests/ut/kv_offload/test_native_cpu_offload.py` | — | Block -> chunk
API: assert `spec.num_chunks` |
| `tests/ut/kv_offload/test_npu_offload_spec.py` | — |
`num_blocks`/`kv_bytes_per_block` -> `num_chunks`/`kv_bytes_per_chunk` |
| `tests/ut/models/test_deepseek_v41_registration.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip; registration is always active |
| `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py` | —
| Removed the 0.29 gate for `mamba_fine_grained_prefix_cache` |
| `tests/ut/patch/platform/test_patch_speculative_config_dspark.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Removed
the DeepSeek V4.1 0.29 `pytest.skip` |
| `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Deleted
the 0.28.0 PCP+DP validator workaround tests; reworked the
version-routing fixtures |
| `tests/ut/patch/worker/test_patch_dspark_pp.py` | — | Parametrize
`legacy: bool` instead of version strings |
| `tests/ut/quantization/configs/test_modelslim_config.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip |
| `tests/ut/spec_decode/test_dspark_proposer.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip marker |
| `tests/ut/spec_decode/test_eagle_proposer.py` | — | `uses_xdrope_dim`
-> `mrope_num_dims` |
| `tests/ut/test_compressed_prefix_cache.py` | — | `replay_boundaries`
now unconditional |
| `tests/ut/worker/test_encoder_acl_graph.py` | — | `axis_keys=()` now
unconditional |
| `tests/ut/worker/test_model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skips and the `vllm_version_is` patches |
| `tests/ut/worker/test_model_runner_v2.py` | v0.30.0 boundary |
Version-routing fixture rework |
| `tests/ut/worker/test_pcp_manager_v2.py` |
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) | Dropped
the 0.29 `req_states` coverage / added `padded_num_reqs` coverage;
removed the 0.28/0.29 branches |
| `tests/ut/worker/v2/test_pp_utils.py` | — | Replaced the
version-routing matrix with `use_legacy_spec_pp() is False` |
| `vllm_ascend/_310p/model_runner_310p.py` | — | Removed the 0.29
xdrope-position branch |
| `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py` | — | Removed
the 0.29 xdrope `target_positions[0]` squeeze |
| `vllm_ascend/_310p/worker/v2/model_runner.py` |
[vllm#57270](vllm-project/vllm#57270) | Removed
the 0.29 `max_seq_len_np` kwarg and the 0.29 eagle-block-drop branch |
| `vllm_ascend/_310p/worker/v2/rope.py` | — | mrope `num_dims =
model_config.mrope_num_dims` (dropped the 0.29 constant 3) |
| `vllm_ascend/_310p/worker/v2/states.py` |
[vllm#56908](vllm-project/vllm#56908) |
`UvaBuffer.uva` property -> method; final v0.30.0 form |
| `vllm_ascend/attention/mla_v1.py` |
[vllm#56181](vllm-project/vllm#56181) | Draft
TND_NTD layout forced via `_EXTRA_CTX.is_draft_model`; dropped the 0.29
gate |
| `vllm_ascend/attention/utils.py` |
[vllm#55353](vllm-project/vllm#55353) /
[vllm#56157](vllm-project/vllm#56157) |
Ascend-owned
`_seq_lens_cpu`/`_num_computed_tokens_cpu`/`dcp_local_seq_lens_cpu` are
now unconditional |
| `vllm_ascend/batch_invariant.py` | — | `reduce_sum` accepts
NumPy-style `axis` + `dtype`, rejects `dim`+`axis` together, forwards
`dtype` to the native fallback |
| `vllm_ascend/core/dyntra_lb_scheduler.py` | — |
`KVConnectorBlockState` always uses `req_ids`+`resolve_block_ids`
(dropped the 0.29 `block_ids` snapshot) |
| `vllm_ascend/core/kv_cache_interface.py` |
[vllm#53906](vllm-project/vllm#53906) | MLA
`get_storage_block_size` override unconditional; dropped the 0.29
`storage_block_size` property |
| `vllm_ascend/core/recompute_scheduler.py` | — | Dropped the xdrope
kwarg and the 0.29 `block_ids` snapshot |
| `vllm_ascend/core/scheduler_profiling_chunk.py` | — | Dropped the 0.29
`block_ids` snapshot |
| `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`
| — | Block -> chunk API (`num_chunks`, `kv_bytes_per_chunk`)
unconditional |
| `vllm_ascend/lora/punica_npu.py` |
[vllm#53555](vllm-project/vllm#53555) + boundary
| `add_lora_logits` per-adapter matmul fallback for heads smaller than
the rank; final state drops the `apply_lora_full_linear` binding (both
supported targets predate #53555) |
| `vllm_ascend/models/__init__.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1
registration (`DeepseekV41ForCausalLM`/`DSparkModel`) unconditional |
| `vllm_ascend/models/deepseek_v41/engram/embedding.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/hash_state.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/parallel.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/model.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/vl_model.py` |
[vllm#56741](vllm-project/vllm#56741) /
[vllm#56554](vllm-project/vllm#56554) |
`deepseek_v41` imports; drop `IMAGE_PAD_ID`/alignment-pad handling on
main |
| `vllm_ascend/ops/mla.py` |
[vllm#56157](vllm-project/vllm#56157) |
`MLAAttention.supports_pcp_dcp = True` set on the class, unconditional |
| `vllm_ascend/ops/rotary_embedding.py` |
[vllm#56446](vllm-project/vllm#56446) | YaRN
mscale signature adaptation; v0.29 branch collapsed |
| `vllm_ascend/patch/__init__.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Registry
entry for the new draft-EP patch; dropped version gates |
| `vllm_ascend/patch/platform/__init__.py` |
[vllm#56741](vllm-project/vllm#56741 Engram
| `patch_engram_config` imported unconditionally |
| `vllm_ascend/patch/platform/patch_balance_schedule.py` | — | Dropped
the 0.29 `block_ids` snapshot |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#54736](vllm-project/vllm#54736) + boundary
| Accept/forward `allow_partial_hash_hits`; collapsed the 0.29 gate |
| `vllm_ascend/patch/platform/patch_parallel_config.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.29 `_validate_parallel_config` PCP+DP workaround; keeps
`use_sequence_parallel_moe` |
| `vllm_ascend/patch/platform/patch_speculative_config.py` |
[vllm#55914](vllm-project/vllm#55914) without
[vllm#56930](vllm-project/vllm#56930) | Skip
`_verify_with_expert_parallelism` for non-MoE draft
(`runner_type=="draft"`) |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.28.0 PCP+DP validation workaround |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py` |
[vllm#53781](vllm-project/vllm#53781) | Ascend
`bind_kv_cache_to_layers` (assign the raw allocation); collapsed the
0.29 gate |
| `vllm_ascend/patch/worker/patch_deepseek_v2.py` |
[vllm#53781](vllm-project/vllm#53781) |
Accept/ignore `index_group_builder`; `SparseMLAIndexGroupBuilder` import
collapse |
| `vllm_ascend/patch/worker/patch_mamba_utils.py` |
[vllm#56898](vllm-project/vllm#56898) |
`GPUInputBatch` import source version-gated, then collapsed to
`gpu_input_batch.InputBatch` |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py` |
[vllm#53781](vllm-project/vllm#53781) | Register
Ascend `bind_kv_cache_to_layers`; expose tuple element 0 for the device
filter; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | — | Legacy
Spec+PP bypass comment (inactive on v0.30.0) |
| `vllm_ascend/patch/worker/patch_v2/patch_spec_pp.py` |
[vllm#56888](vllm-project/vllm#56888) | Alias
`async_tensor_h2d as async_copy_to_gpu`; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_uva.py` |
[vllm#56908](vllm-project/vllm#56908) | `uva`
property vs method; final v0.30.0 method form |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#56254](vllm-project/vllm#56254) +
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) + boundary
| Gate the V4.1 DSpark import by `HAS_TRITON`; static
`_draft_embed_accepts_mm` check instead of the runtime `embed_input_ids`
probe; collapsed version gates |
| `vllm_ascend/utils.py` | — | `vllm_version_is` docstring 0.29 -> 0.30;
Kimi MLA custom-op registration unconditional |
| `vllm_ascend/worker/model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1 dsa
metadata/cache imports unconditional; removed xdrope position handling |
| `vllm_ascend/worker/v2/aclgraph_utils.py` |
[vllm#51700](vllm-project/vllm#51700) |
`ModelAclGraphManager.__init__` accepts/forwards `ubatch_runner`;
`UBatchRunner` import collapse |
| `vllm_ascend/worker/v2/model_runner.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#53867](vllm-project/vllm#53867) /
[vllm#51700](vllm-project/vllm#51700) /
[vllm#57270](vllm-project/vllm#57270) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async-copy alias; pass `BatchExecutionDescriptor` to
`maybe_partition_pcp_batch`; `ubatch_runner`; `make_dummy(is_padding)`;
keep the replicated PCP draft on the global batch; collapsed version
gates. Also restores the `_check_oproj_tp_graph_step` guard and the PD
decode-recompute `gather_batch_req_state` override accidentally removed
by the v0.28.0-boundary cleanup (`e30adae80`) |
| `vllm_ascend/worker/v2/pcp_manager.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async alias; drop `req_states`; forward `padded_num_reqs`;
`prepare_draft_prefill` no-op / `restore_for_sampling` skip; collapsed
version gates |
| `vllm_ascend/worker/v2/pp_utils.py` | — | `use_legacy_spec_pp()`
returns `False` |
| `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Route
through `_build_uniform_attn_metadata`/`_build_attn_metadata`; re-hook
the Ascend rotary-positions injection onto the new methods; dropped the
0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
`BatchExecutionDescriptor` routing; dropped the 0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
routing; dropped the 0.29 gate |

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
hw-qianlingfeng pushed a commit to hw-qianlingfeng/vllm-ascend that referenced this pull request Sep 27, 2026
### What this PR does / why we need it?

#### Change Summary

The PR advances the main2main lane to vLLM v0.30.0 (commit
`ced6857afa0ea7b2e3f0846a62e1394e90f15607`), adapting vllm-ascend to
every upstream change in the `84030bbe` -> `4991f97` -> `ced6857` range.
Because both CI lanes now install vLLM v0.30.0, all
`vllm_version_is("0.29.0")` forks are permanently false and are
collapsed to the main behavior.

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
| `.github/vllm-main-verified.commit` | — | Updated verified main commit
hash `84030bbe` -> `4991f97` -> `ced6857` (v0.30.0 tag commit) |
| `.github/vllm-release-tag.commit` | — | Bumped the release boundary to
`v0.30.0` |
| `.github/workflows/pr_test.yaml` | — | Added a dual-version cpu-ut
matrix (`vllm_versions`) for the main2main lane; temporarily commented
out pre-commit/mypy and set `cpu-ut` to `if: false`; dropped the
`needs.cpu-ut` requirement from the ready gate; added a
main2main-specific cpu-ut failure hint |
| `Dockerfile` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.310p.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a3.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.a5.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `Dockerfile.openEuler` | — | `VLLM_TAG` default -> `v0.30.0` |
| `README.md` | — | CI notes updated to v0.30.0 |
| `README.zh.md` | — | CI notes updated to v0.30.0 |
|
`tests/e2e/nightly/single_node/ops/singlecard_ops/triton/test_num_nans.py`
| — | `VLLM_VERSION` default 0.29.0 -> 0.30.0 |
| `tests/ut/_310p/test_model_runner_v2_310p.py` | v0.30.0 boundary |
Dropped the `vllm_version_is` gate test; `_needs_kv_cache_zeroing_310p`
always uses `spec_config.use_eagle_block_drop()` |
| `tests/ut/core/test_dyntra_lb_scheduler.py` | — | Removed the 0.29
`KVConnectorBlockState.block_ids` assertions; assert the `req_ids` form
only |
| `tests/ut/core/test_scheduler_connector_block_state.py` | — | Removed
the 0.29 `block_ids` snapshot branch |
| `tests/ut/kv_offload/test_native_cpu_offload.py` | — | Block -> chunk
API: assert `spec.num_chunks` |
| `tests/ut/kv_offload/test_npu_offload_spec.py` | — |
`num_blocks`/`kv_bytes_per_block` -> `num_chunks`/`kv_bytes_per_chunk` |
| `tests/ut/models/test_deepseek_v41_registration.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip; registration is always active |
| `tests/ut/patch/platform/test_patch_mamba_block_aligned_split.py` | —
| Removed the 0.29 gate for `mamba_fine_grained_prefix_cache` |
| `tests/ut/patch/platform/test_patch_speculative_config_dspark.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Removed
the DeepSeek V4.1 0.29 `pytest.skip` |
| `tests/ut/patch/platform/test_patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Deleted
the 0.28.0 PCP+DP validator workaround tests; reworked the
version-routing fixtures |
| `tests/ut/patch/worker/test_patch_dspark_pp.py` | — | Parametrize
`legacy: bool` instead of version strings |
| `tests/ut/quantization/configs/test_modelslim_config.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip |
| `tests/ut/spec_decode/test_dspark_proposer.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skip marker |
| `tests/ut/spec_decode/test_eagle_proposer.py` | — | `uses_xdrope_dim`
-> `mrope_num_dims` |
| `tests/ut/test_compressed_prefix_cache.py` | — | `replay_boundaries`
now unconditional |
| `tests/ut/worker/test_encoder_acl_graph.py` | — | `axis_keys=()` now
unconditional |
| `tests/ut/worker/test_model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | Removed
the V4.1 0.29 skips and the `vllm_version_is` patches |
| `tests/ut/worker/test_model_runner_v2.py` | v0.30.0 boundary |
Version-routing fixture rework |
| `tests/ut/worker/test_pcp_manager_v2.py` |
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) | Dropped
the 0.29 `req_states` coverage / added `padded_num_reqs` coverage;
removed the 0.28/0.29 branches |
| `tests/ut/worker/v2/test_pp_utils.py` | — | Replaced the
version-routing matrix with `use_legacy_spec_pp() is False` |
| `vllm_ascend/_310p/model_runner_310p.py` | — | Removed the 0.29
xdrope-position branch |
| `vllm_ascend/_310p/spec_decode/llm_base_proposer_310.py` | — | Removed
the 0.29 xdrope `target_positions[0]` squeeze |
| `vllm_ascend/_310p/worker/v2/model_runner.py` |
[vllm#57270](vllm-project/vllm#57270) | Removed
the 0.29 `max_seq_len_np` kwarg and the 0.29 eagle-block-drop branch |
| `vllm_ascend/_310p/worker/v2/rope.py` | — | mrope `num_dims =
model_config.mrope_num_dims` (dropped the 0.29 constant 3) |
| `vllm_ascend/_310p/worker/v2/states.py` |
[vllm#56908](vllm-project/vllm#56908) |
`UvaBuffer.uva` property -> method; final v0.30.0 form |
| `vllm_ascend/attention/mla_v1.py` |
[vllm#56181](vllm-project/vllm#56181) | Draft
TND_NTD layout forced via `_EXTRA_CTX.is_draft_model`; dropped the 0.29
gate |
| `vllm_ascend/attention/utils.py` |
[vllm#55353](vllm-project/vllm#55353) /
[vllm#56157](vllm-project/vllm#56157) |
Ascend-owned
`_seq_lens_cpu`/`_num_computed_tokens_cpu`/`dcp_local_seq_lens_cpu` are
now unconditional |
| `vllm_ascend/batch_invariant.py` | — | `reduce_sum` accepts
NumPy-style `axis` + `dtype`, rejects `dim`+`axis` together, forwards
`dtype` to the native fallback |
| `vllm_ascend/core/dyntra_lb_scheduler.py` | — |
`KVConnectorBlockState` always uses `req_ids`+`resolve_block_ids`
(dropped the 0.29 `block_ids` snapshot) |
| `vllm_ascend/core/kv_cache_interface.py` |
[vllm#53906](vllm-project/vllm#53906) | MLA
`get_storage_block_size` override unconditional; dropped the 0.29
`storage_block_size` property |
| `vllm_ascend/core/recompute_scheduler.py` | — | Dropped the xdrope
kwarg and the 0.29 `block_ids` snapshot |
| `vllm_ascend/core/scheduler_profiling_chunk.py` | — | Dropped the 0.29
`block_ids` snapshot |
| `vllm_ascend/distributed/kv_transfer/kv_pool/kv_offload/native/npu.py`
| — | Block -> chunk API (`num_chunks`, `kv_bytes_per_chunk`)
unconditional |
| `vllm_ascend/lora/punica_npu.py` |
[vllm#53555](vllm-project/vllm#53555) + boundary
| `add_lora_logits` per-adapter matmul fallback for heads smaller than
the rank; final state drops the `apply_lora_full_linear` binding (both
supported targets predate #53555) |
| `vllm_ascend/models/__init__.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1
registration (`DeepseekV41ForCausalLM`/`DSparkModel`) unconditional |
| `vllm_ascend/models/deepseek_v41/engram/embedding.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/hash_state.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/engram/parallel.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/model.py` |
[vllm#56741](vllm-project/vllm#56741) |
`deepseek_v4_1` -> `deepseek_v41` import unconditional |
| `vllm_ascend/models/deepseek_v41/vl_model.py` |
[vllm#56741](vllm-project/vllm#56741) /
[vllm#56554](vllm-project/vllm#56554) |
`deepseek_v41` imports; drop `IMAGE_PAD_ID`/alignment-pad handling on
main |
| `vllm_ascend/ops/mla.py` |
[vllm#56157](vllm-project/vllm#56157) |
`MLAAttention.supports_pcp_dcp = True` set on the class, unconditional |
| `vllm_ascend/ops/rotary_embedding.py` |
[vllm#56446](vllm-project/vllm#56446) | YaRN
mscale signature adaptation; v0.29 branch collapsed |
| `vllm_ascend/patch/__init__.py` |
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) | Registry
entry for the new draft-EP patch; dropped version gates |
| `vllm_ascend/patch/platform/__init__.py` |
[vllm#56741](vllm-project/vllm#56741 Engram
| `patch_engram_config` imported unconditionally |
| `vllm_ascend/patch/platform/patch_balance_schedule.py` | — | Dropped
the 0.29 `block_ids` snapshot |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#54736](vllm-project/vllm#54736) + boundary
| Accept/forward `allow_partial_hash_hits`; collapsed the 0.29 gate |
| `vllm_ascend/patch/platform/patch_parallel_config.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.29 `_validate_parallel_config` PCP+DP workaround; keeps
`use_sequence_parallel_moe` |
| `vllm_ascend/patch/platform/patch_speculative_config.py` |
[vllm#55914](vllm-project/vllm#55914) without
[vllm#56930](vllm-project/vllm#56930) | Skip
`_verify_with_expert_parallelism` for non-MoE draft
(`runner_type=="draft"`) |
| `vllm_ascend/patch/platform/patch_use_v2_model_runner.py` |
[vllm#54523](vllm-project/vllm#54523) | Removed
the 0.28.0 PCP+DP validation workaround |
| `vllm_ascend/patch/worker/patch_bind_kv_cache.py` |
[vllm#53781](vllm-project/vllm#53781) | Ascend
`bind_kv_cache_to_layers` (assign the raw allocation); collapsed the
0.29 gate |
| `vllm_ascend/patch/worker/patch_deepseek_v2.py` |
[vllm#53781](vllm-project/vllm#53781) |
Accept/ignore `index_group_builder`; `SparseMLAIndexGroupBuilder` import
collapse |
| `vllm_ascend/patch/worker/patch_mamba_utils.py` |
[vllm#56898](vllm-project/vllm#56898) |
`GPUInputBatch` import source version-gated, then collapsed to
`gpu_input_batch.InputBatch` |
| `vllm_ascend/patch/worker/patch_v2/patch_attn_utils.py` |
[vllm#53781](vllm-project/vllm#53781) | Register
Ascend `bind_kv_cache_to_layers`; expose tuple element 0 for the device
filter; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_dspark.py` | — | Legacy
Spec+PP bypass comment (inactive on v0.30.0) |
| `vllm_ascend/patch/worker/patch_v2/patch_spec_pp.py` |
[vllm#56888](vllm-project/vllm#56888) | Alias
`async_tensor_h2d as async_copy_to_gpu`; collapsed the 0.29 gate |
| `vllm_ascend/patch/worker/patch_v2/patch_uva.py` |
[vllm#56908](vllm-project/vllm#56908) | `uva`
property vs method; final v0.30.0 method form |
| `vllm_ascend/spec_decode/llm_base_proposer.py` |
[vllm#56254](vllm-project/vllm#56254) +
[vllm#55914](vllm-project/vllm#55914) /
[vllm#56930](vllm-project/vllm#56930) + boundary
| Gate the V4.1 DSpark import by `HAS_TRITON`; static
`_draft_embed_accepts_mm` check instead of the runtime `embed_input_ids`
probe; collapsed version gates |
| `vllm_ascend/utils.py` | — | `vllm_version_is` docstring 0.29 -> 0.30;
Kimi MLA custom-op registration unconditional |
| `vllm_ascend/worker/model_runner_v1.py` |
[vllm#56741](vllm-project/vllm#56741) | V4.1 dsa
metadata/cache imports unconditional; removed xdrope position handling |
| `vllm_ascend/worker/v2/aclgraph_utils.py` |
[vllm#51700](vllm-project/vllm#51700) |
`ModelAclGraphManager.__init__` accepts/forwards `ubatch_runner`;
`UBatchRunner` import collapse |
| `vllm_ascend/worker/v2/model_runner.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#53867](vllm-project/vllm#53867) /
[vllm#51700](vllm-project/vllm#51700) /
[vllm#57270](vllm-project/vllm#57270) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async-copy alias; pass `BatchExecutionDescriptor` to
`maybe_partition_pcp_batch`; `ubatch_runner`; `make_dummy(is_padding)`;
keep the replicated PCP draft on the global batch; collapsed version
gates. Also restores the `_check_oproj_tp_graph_step` guard and the PD
decode-recompute `gather_batch_req_state` override accidentally removed
by the v0.28.0-boundary cleanup (`e30adae80`) |
| `vllm_ascend/worker/v2/pcp_manager.py` |
[vllm#56888](vllm-project/vllm#56888) /
[vllm#56107](vllm-project/vllm#56107) /
[vllm#53867](vllm-project/vllm#53867) +
[vllm#56181](vllm-project/vllm#56181) + boundary
| Async alias; drop `req_states`; forward `padded_num_reqs`;
`prepare_draft_prefill` no-op / `restore_for_sampling` skip; collapsed
version gates |
| `vllm_ascend/worker/v2/pp_utils.py` | — | `use_legacy_spec_pp()`
returns `False` |
| `vllm_ascend/worker/v2/spec_decode/autoregressive/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Route
through `_build_uniform_attn_metadata`/`_build_attn_metadata`; re-hook
the Ascend rotary-positions injection onto the new methods; dropped the
0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dflash/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
`BatchExecutionDescriptor` routing; dropped the 0.29 gate |
| `vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` |
[vllm#56181](vllm-project/vllm#56181) | Same
routing; dropped the 0.29 gate |

### Does this PR introduce _any_ user-facing change?

### How was this patch tested?


- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

mrv2 Model Runner V2 specific nvidia ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm

Projects

Status: Done
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants